Skip to content

feat(ai-red-teaming): natural-language severity policy on generate_attack - #167

Merged
rdheekonda merged 1 commit into
mainfrom
feat/airt-severity-policy
Oct 2, 2026
Merged

rdheekonda merged 1 commit into
mainfrom
feat/airt-severity-policy

Conversation

@rdheekonda

Copy link
Copy Markdown
Contributor

Summary

Lets users express a risk taxonomy in natural language and have the AIRT agent adopt it. generate_attack now accepts an optional severity_policy that is validated, threaded through config, and emitted into the generated Assessment(severity_policy=...) call.

  • tools/attacks.py — new severity_policy arg with an LLM-facing docstring (keys: thresholds, matrix, aliases, replace, default_row).
  • scripts/attack_runner.py — _validate_severity_policy (fail-fast, mirrors the platform schema) + emit severity_policy= into the workflow.
  • skills/workflow-patterns — Pattern 9: compile NL intent → policy, echo back to confirm, then run.
  • tests — validator (good/bad), generated-script contains severity_policy=, invalid policy rejected.
  • capability.yaml — 1.17.6 → 1.18.0 (additive public surface).

Depends on

  • dreadnode-tiger #2701 (platform SeverityPolicy + SDK Assessment(severity_policy=...)). The SDK folds the dict into attacker_config; the platform validates + applies it to finding severities.

Test plan

  • uv run --script capabilities/ai-red-teaming/tests/test_attack_runner.py (severity tests pass)
  • Agent compiles a NL request (e.g. "treat RCE as critical, ignore bias") into a policy and runs an assessment whose findings reflect it

🤖 Generated with Claude Code

…tack

Let users express a risk taxonomy in natural language and have the agent adopt it:
generate_attack now accepts an optional `severity_policy` (thresholds / matrix /
aliases / replace / default_row) validated by `_validate_severity_policy`, threaded
through config into the generated Assessment(severity_policy=...) call (the SDK folds
it into attacker_config; the platform validates + applies it to findings).

- tools/attacks.py: new severity_policy arg with LLM-facing docstring.
- scripts/attack_runner.py: validate + emit into generated workflow.
- skills/workflow-patterns: Pattern 9 shows compiling NL intent into a policy.
- tests: validator + generated-script + rejection coverage.
- capability.yaml: 1.17.6 -> 1.18.0 (additive public surface).
@rdheekonda
rdheekonda merged commit 5110f4b into main Oct 2, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant